Back

Cell Systems

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Cell Systems's content profile, based on 201 papers previously published here. The average preprint has a 0.20% match score for this journal, so anything above that is already an above-average fit.

1
Impacts of batch effects on the performance of machine learning classifiers across multiple studies

Raab, P.; Johnson, W. E.; Piccolo, S. R.

2026-06-30 bioinformatics 10.64898/2026.06.24.734352 medRxiv
Top 0.1%
39.9%
Show abstract

Precision medicine relies on accurate and generalizable predictions for patients across the spectrum of human diversity. Because capturing biological heterogeneity requires large sample sizes, researchers must often aggregate data from several experimental batches or independent studies. This integration allows for greater statistical power and diversity than a single study could provide, while avoiding the costs of generating massive new -omics datasets. Predictive models trained on these aggregated data are theoretically better equipped to detect subtle patterns that generalize to new data. However, this potential is frequently undermined by "batch effects"--systematic technical artifacts that can bias model training to predict experimental batches and shadow meaningful biological conditions. Models trained on data with batch effects can exhibit substantially degraded performance when applied to data from new batches. Statistical adjustment methods can mitigate these artifacts while preserving biological signals. To ensure these adjustments actually facilitate generalization, we emphasize the use of external, independent cohorts for rigorous validation. This chapter examines how batch effects impact predictions and compares various adjustment methods.

2
Breaking the Synthesis Barrier for AI-Designed DNA Libraries

Sussex, S.; Borevkovic, E.; Lohmann, F.; Chen, N.; Lüthi, E.; Reddy, S. T.; Krause, A.

2026-07-07 bioengineering 10.64898/2026.07.07.736931 medRxiv
Top 0.1%
38.3%
Show abstract

Designing DNA libraries is a key challenge from drug design to protein engineering and synthetic biology. Modern generative models offer opportunities to navigate the design space and propose specific sequences predicted to be effective in-silico. Designing deterministic libraries of specific sequences is however limited by the cost of DNA synthesis -- the synthesis barrier. In contrast, high-throughput multiplexed screening can measure the function of billions of biological sequences in parallel. Harnessing this technology requires the design of randomized libraries with specific design constraints to achieve low synthesis costs. In practice, such stochastic libraries are often chosen heuristically, sacrificing control for scale. Is there a way to bridge AI-based in-silico sequence design with high-throughput experimentation? In this work, we introduce Policy Gradients for Library Design (PGLD). PGLD uses a synthesis-aware parametrization of stochastic DNA libraries and optimizes them against a specified objective function. This allows for designing massive, controlled libraries without being limited by synthesis costs. We show how PGLD enables lab-in-the-loop design of multi-round high-throughput experiments, and large-scale in-vitro DNA sampling from generative models. Finally, we use PGLD to design a library of ~10^6 unique sequences which is synthesized at a cost of ~700 USD to explore the mutation space of a broadly neutralizing influenza antibody.

3
Task-adapted biological foundation models uncover perturbation-centric representations

Pareja-Lorente, E.; Aloy, P.

2026-07-05 bioinformatics 10.64898/2026.06.30.735584 medRxiv
Top 0.1%
23.1%
Show abstract

Foundation models have emerged as powerful tools for learning transferable representations of biological systems, yet their latent spaces are typically optimized to capture cellular state rather than the effects of perturbations. Here, we demonstrate that a biological foundation model can be repurposed to learn a fundamentally different representation by changing its learning objective. We finetuned scGPT, a transformer pre-trained on over 30 million single cell transcriptomes, on more than three million LINCS L1000 perturbation profiles using a supervised objective that predicts perturbation identity. This transformed the latent space into a perturbation centric representation that aligned transcriptional responses induced by the same chemical or genetic perturbation across heterogeneous experimental conditions. Finetuned embeddings substantially outperformed both gene expression profiles and the original pretrained model, recovering [85, 100%] of perturbations within the top 100 nearest neighbors and increasing perturbation classification accuracy from [10, 19%] to [25, 49%]. Remarkably, although the model was trained exclusively to recognize perturbation identity, the learned representation spontaneously captured orthogonal biological relationships never provided during training, including chemical similarity (AUROC up to 0.81), mechanisms of action (Hit@10 up to 100%), compound target relationships (AUROC up to 0.74), and functional relationships between genetic perturbations. The resulting embedding space enabled mechanism-of-action annotation of nearly 12,000 previously uncharacterized compounds, prioritization of target related chemical genetic associations, and contextualization of unseen perturbations and external transcriptomic datasets. Together, our results establish objective-driven adaptation as a general strategy for repurposing biological foundation models to learn reusable representations of complex biological phenomena.

4
Unbalanced Perturbation Dynamics For Cell Fate Design

Peng, Q.; Wang, Y.; Li, J.; Wang, X.; Xiao, Y.; Zhou, P.

2026-07-04 bioinformatics 10.64898/2026.06.30.735555 medRxiv
Top 0.1%
22.2%
Show abstract

Large-scale single-cell perturbation sequencing provides an unprecedented opportunity to construct virtual cells for the in silico simulation of cellular responses and the inverse design of optimal interventions. However, most perturbation-response models treat cellular responses primarily as mass-preserving shifts in transcriptomic state, whereas single-cell perturbation measurements are inherently unbalanced: the recovered endpoint population is shaped by technical sampling as well as biological perturbation-induced proliferation, apoptosis and selection. Here we introduce U-Pert, an unbalanced generative framework that learns condition- and context-dependent perturbation dynamics from unpaired single-cell snapshots. U-Pert jointly models transcriptomic state transitions and cell-number dynamics, enabling scalable and robust forward prediction of unseen perturbations and contexts, as well as inverse design to screen for desired genetic or pharmacological interventions that achieve user-defined transcriptomic or population-level outcomes. Across controlled simulations, genetic perturbation benchmarks, sciPlex3 drug responses and PBMC cytokine perturbations, U-Pert predicts unseen responses, captures both molecular and abundance changes, and performs inverse design for target gene-expression programs and cell-type compositions. These results show that cell abundance is an integral component of the perturbation phenotype, providing a mass-aware framework for virtual-cell modeling and perturbation cell fate design.

5
Reliability-weighted target prioritization in CD4+ T-cell Perturb-seq: a generalizability-theory decomposition

Cheng, C.

2026-07-15 bioinformatics 10.64898/2026.07.13.738312 medRxiv
Top 0.1%
21.9%
Show abstract

Genome-scale Perturb-seq screens prioritize candidate targets by the strength of a perturbations transcriptional effect. Effect strength does not answer a prior measurement question: is the readout dependable? A large effect estimated from a single guide, a single donor, or a pseudobulk of few cells need not survive replication, and for target prioritization each false lead costs a validation experiment. We treat each perturbation effect as a measurement in a crossed Target x Guide x Donor x Condition design and apply generalizability theory (Brennan, 2001; Cronbach et al., 1972) to separate the dependable part of an effect from facet-specific idiosyncrasy. Guides and donors enter as random facets; condition enters as a fixed facet and is analyzed within its levels. For each target we report a dependability profile over the facets and a joint generalizability coefficient over the two random facets, and we re-rank targets by effect magnitude weighted by that coefficient. On the released screen (Zhu et al., 2025), removing the measurement-error floor estimated from the non-targeting controls raises the number of genes with a dependable target-signal share above .10 from 40 to 7,674. Analyzed within activation states, dependability recovers the T-cell-receptor signaling module as reliably measurable only in activated cells, without recourse to gene annotation. A design study indicates that reliability is limited by the number of guides rather than the number of donors, so a future screen should add guides. Every methodological decision was recorded and adversarially reviewed, and all results regenerate from the released summary statistics.

6
ORBIT: Annotation-Aware Empirical Enrichment and Semantic Reranking for Interpretable Functional-Class Recovery

Kidder, B. L.

2026-07-07 bioinformatics 10.64898/2026.07.01.735870 medRxiv
Top 0.1%
21.8%
Show abstract

Gene-set interpretation workflows are widely used to summarize transcriptomic and proteomic experiments, yet standard enrichment tools often return long, redundant result tables that require substantial manual consolidation. We developed ORBIT (Ontology-Ranked Biological Interpretation Tool), an annotation-aware interpretation workflow that combines empirical enrichment, semantic reranking, and redundancy-aware representative-term selection to prioritize interpretable functional summaries from gene sets. We evaluated ORBIT on a curated tiered benchmark of human functional-class gene sets spanning clean reference sets, size-ladder variants, and mixed-difficulty cases. On the 45-set core benchmark, ORBIT semantic achieved higher expected-class recovery than Enrichr and PANTHER Gene Ontology molecular-function baselines, with a mean reciprocal rank of 0.916 and top-1 recovery of 0.889. Bootstrap confidence intervals and paired permutation testing supported the robustness of this advantage, and supplemental analyses extended the comparison to g:Profiler. In a GPCR mixed-function case study, ORBIT compressed redundant enriched terms into semantic representative neighborhoods, illustrating how long enrichment outputs can be converted into reviewable biological summaries. We then used ORBIT to interpret immune-cell identity, interferon-response biology, and breast-cancer subtype programs. ORBIT linked PBMC3K markers to cytotoxic, antigen-presentation, and innate-immune cell states; prioritized antiviral, cytokine-response, RNA-binding, and secreted-factor biology after IFNB stimulation; and separated TCGA-BRCA basal-like proliferative chromosome/cell-cycle programs from luminal transporter and receptor-associated biology while retaining gene-level support.

7
The structural context of mutations in proteins predicts their effect on antibiotic resistance

Green, A. G.; Tasmin, M.; Vargas, R.; Farhat, M. R.

2026-06-29 bioinformatics 10.1101/2025.09.23.676583 medRxiv
Top 0.1%
18.9%
Show abstract

In Mycobacterium tuberculosis, a prevalent and deadly pathogen, resistance to antibiotics evolves primarily through non-synonymous mutations in proteins. Sequence-based analyses are currently used to understand the genetic basis of antibiotic resistance, either via genotype-phenotype association, or via signals of convergent evolution. These methods focus on primary sequence and often neglect other biological signals such as protein structural information. We hypothesize that integrating the structural context of mutations improves the prediction of effects on function and phenotype. We curate high confidence structural annotations for the M. tuberculosis proteome from 1,371 crystallography and 2,316 AlphaFold predictions, and combine the structures with mutations from over 31,000 clinical M. tuberculosis isolates. We demonstrate that mutations in proteins known to cause resistance are clustered in 3D space, even in proteins where inactivating mutations at any position are thought to cause resistance. We develop a statistic to search the M. tuberculosis proteome for signal of clustered mutations, finding over 450 proteins that display this signal, many of which have a known relationship with antibiotic resistance. We show that a supervised classifier trained on 3D distance to known resistance sites alone has an F1 score of 94.6% at classifying mutations as resistance-conferring across proteins. This work demonstrates that protein structure provides useful information for categorizing which variants may cause antibiotic resistance, even when the majority of structures are AI-predicted.

8
V3Cell: A Vision-Guided Virtual 3D Cell Framework for Phenotypic Modeling and Perturbation Prediction

Lu, Y.; Xun, D.; chenke, X.; Xiaobo, Z.; Zhigang, Z.; Pengyu, C.; Xiwen, Y.; Zhengzheng, Y.; Jiahua, R.; Huili, H.; Jianying, H.; Pengwei, H.

2026-06-24 bioinformatics 10.64898/2026.06.23.734130 medRxiv
Top 0.1%
18.9%
Show abstract

Predicting how organoids respond to chemical perturbations is central to disease modeling and drug discovery. Existing virtual cell models operate at the single-cell level, producing static endpoint predictions from destructive assays. This leaves a critical gap at the organoid scale, where biological identity is defined by tissue-level architecture and continuous developmental dynamics rather than single-cell features. Here we introduce V3Cell, a vision-guided framework that constructs in silico surrogates of organoids directly from non-invasive bright-field microscopy. A foreground-aware model constructs static virtual 3D cells across colon, stomach, and lung organoid lineages. These virtual 3D cells closely match real samples across distributional metrics, micro-texture, and lineage-specific morphometrics, with small effect sizes for most descriptors. A temporal module further predicts developmental fate from as few as six early-frame observations and models fate-conditioned spatiotemporal trajectories that closely recapitulate real perturbation responses. V3Cell requires no omics profiling or fluorescent labeling, establishing a non-invasive brightfield-based paradigm for organoid-scale perturbation prediction. Our code and data are publicly available at https://github.com/Laineyoulu/V3Cell.

9
DELPHAI predicts heterogeneous perturbation responses with learned cell fitness and gene-space retrieval

Zhang, X.; Wu, H.; Liu, H.

2026-07-06 bioinformatics 10.64898/2026.07.01.735965 medRxiv
Top 0.1%
18.6%
Show abstract

Modelling heterogeneous cellular responses to perturbation holds the promise of scalable in silico screening and mechanistic insight. However, mass conservation despite cell-type-specific depletion, and lossy projections from gene space to latent space, hinder performance of state-of-the-art methods. DELPHAI, with learned per-cell-fitness filtering out depleted cells and gene-space retrieval bypassing the latent bottleneck, outperforms all baseline methods across two benchmark frameworks and offers explainability with inferred cell-type-specific survival without any biological priors.

10
SPARC: A Graph-based Optimization Framework for Directional Trajectory Reconstruction Across Ordered Single-Cell Conditions

Wu, S.; Walker, W. C.; Martin, C.; Yustein, J. T.; Samee, M. A. H.

2026-07-11 bioinformatics 10.64898/2026.07.07.736532 medRxiv
Top 0.1%
18.5%
Show abstract

Single-cell transcriptomics has enabled systematic profiling of cellular states across ordered biological contexts, including developmental stages, treatment phases, disease progression, and anatomical compartments. A central challenge is to reconstruct trajectories that respect the directionality imposed by biology or experimental design. Existing trajectory inference methods reconstruct cell-state progressions from latent-space geometry but do not enforce external biological ordering during graph construction, yielding biologically inadmissible transitions. An emerging paradigm of optimal-transport (OT) approaches partially addresses this limitation by incorporating experimental ordering into probabilistic state-to-state correspondences, yet their pairwise formulation cannot resolve whether a given state is an intermediate state or a terminal state along a multi-step progression. In multi-timepoint settings, OT typically estimates couplings only betweenadjacent timepoints and then chains these locally solved couplings to approximate long-range trajectories without a global optimization across all conditions simultaneously. Here we present SPARC, a graph-based optimization framework that quantifies similarity in a shared high-dimensional latent space and reconstruct directional trajectories under biological constraints. Global shortest-path optimization over this graph yields progression routes, from which SPARC derives path-based pseudotime identifies bottlenecks clusters, and detects gene temporal behavior. SPARC was evaluated across three complementary settings representing distinct trajectory-inference challenges. Its application to paired primary and lung metastatic osteosarcoma samples allows us to be the first to propose a "cross-organ bone-like microenvironment" hypothesis, in which osteoclastogenic signaling establishes a bone-like remodeling niche within the pulmonary metastatic lesion that promotes osteoclast differentiation and activity. The findings are independently recoverable in human osteosarcoma Visium HD spatial transcriptomics.

11
The Boolean Breast-Cancer Network (BBCN): durable, controllable apoptosis from faithful feedback and multi-timescale control

Bhatti, A. I.

2026-07-08 systems biology 10.64898/2026.06.18.733086 medRxiv
Top 0.1%
18.5%
Show abstract

We model a breast tumour as a Boolean network of signalling pathways and pose therapy as a control problem: find minimal, druggable interventions that drive the network to a desired cell-fate phenotype as a genuine fixed point of the free dynamics, without permanently forcing any node. A 135-node breast-cancer network, adapted from a published Boolean model and carrying the feedback simplifications usual to such models, driven by a three-tier capped controller, makes the three cell-fate phenotypes controllable to differing degrees on their own (apoptosis 81-86% and proliferation 94-95% of patients across three cohorts, but resistance-off only 4-8%), yet achieving all three in sequence while preserving earlier gains is rare (2-3%), and the baseline network supports no durable apoptotic fixed point at all: the phenotypes share a dominant strongly-connected core (16 of 22 regulatory pathways), and under synchronous updating the network is globally oscillatory. We ask whether this non-durability reflects the biology or the simplified wiring, and show it is the wiring, not the biology. The baseline rules carry a degenerate p53 arm in which MDM2 self-locks and damage never reaches p53, and an AKT-FOXO arm with no closing feedback. Restoring three pieces of textbook regulation, the same biology a reduced bistable switch already uses, a p53-MDM2-ATM damage sensor, an AKT1-FOXO3a loop closed by PHLPP (the 136th node) with an uncoupled-state flag, and a death-engaged commitment latch, changes the picture across three cohorts (TCGA-BRCA, METABRIC, I-SPY2; N=1082/1980/988). Two changes recover durable control and they do different jobs: restoring the three feedback loops to their faithful biological form makes the durable death state exist, while embedding the biological separation of timescales, as a delayed (autoregressive) multirate system with distinct fast-signalling, protein-turnover, and transcription rates, makes that state reachable and stable where a single synchronous clock only oscillates. To our knowledge this direct embedding of timescale separation in the control of a patient-specific Boolean cancer model is novel. Drug resistance becomes a precise molecular state: it equals AKT-FOXO3a uncoupling, a documented resistance mechanism, for 4049 of 4050 patients (99.98%). Survival-axis inhibition durably flips exactly the coupled fraction (33-36%) while genotoxic input is the weak lever (3-4%), reversing the baseline-network reading. A minimal druggable kernel designed on the tractable switch, concentrating on the PI3K/AKT/mTOR axis, is found for 99-100% of resistant patients and commits 92-95% of them to apoptosis on the full network once death is scored by caspase commitment rather than by sustained survival-signal suppression. Under the full staged controller the biology-faithful network supports durable apoptotic fixed points in 15-17% of patients, where the baseline network supported none. We keep two honest readouts of death throughout, a strict one (8-14%) and a commitment one (92-95%), and report both. The model is a structural stratifier and drug-target nominator; it does not predict pathologic complete response, a limitation we trace to endpoint distance rather than signal quality. A minimal switch kernel of about two to three nodes holds this apoptotic state about as durably as a near whole-network controller of about thirty, a roughly 12-fold reduction at equal durability.

12
Systematic benchmarking of zero-shot utility and robustness in single-cell transcriptomic foundation models

Liu, T.; Feng, T.; Pan, X.; Chen, Y.; Ren, L.; Ye, X.; Sakurai, T.; Lin, H.; Zhang, Y.

2026-06-23 bioinformatics 10.64898/2026.06.18.733285 medRxiv
Top 0.1%
18.2%
Show abstract

Single-cell foundation models (scFMs) have been proposed as reusable representations for transcriptomic analysis, yet their practical utility and robustness when applied without task-specific fine-tuning remain incompletely characterized. Here, we systematically evaluated single-cell transcriptomic representations in zero-shot settings across 20 methods, 6 downstream tasks and 1,607 datasets comprising nearly 21.8 million cells. We characterized model behavior along three complementary dimensions: baseline utility, structural robustness, and dataset-level drivers of performance variability. Our large-scale analysis reveals a decoupling between utility and robustness: methods ranking highly on standard benchmarks often show marked instability under shifts in dataset structure. Furthermore, no single model performs uniformly well across tasks. In several tasks, classical statistical representations based on highly variable genes remain competitive under zero-shot conditions. Together, these results define the practical boundaries of zero-shot use in scFMs and provide a large-scale benchmark and decision framework for representation selection in single-cell genomics.

13
Pathogen context reshapes antimicrobial peptide generation

You, S.; Zhang, C.; Han, Y.; Jiang, Q.; Guo, X.; Li, M.; Su, Y.; Dong, X.; Yang, M.; Lu, H.

2026-07-03 bioinformatics 10.64898/2026.07.01.735178 medRxiv
Top 0.1%
18.2%
Show abstract

Antimicrobial peptide discovery is constrained less by the number of molecules that can be generated than by the choice of which few should be tested against a defined pathogen. Most peptide generators produce broadly antimicrobial-like sequences and leave target specificity to downstream filters. Here we show that pathogen context can be introduced during generation. AMPHORA conditions a peptide-native generator on target class, pathogen genome features and strain-description text. Matched, ablated and shuffled controls showed that aligned pathogen inputs redirected generated libraries beyond coarse activity labels, whereas global shuffling weakened this effect. Same-noise counterfactuals showed that strain descriptions drove larger sequence changes, whereas genome features more strongly affected predicted structural properties. Species-level analyses revealed target-dependent enrichment. Matched bacterial inputs also shifted APEX-predicted activity rankings relative to class-only generation. The resulting libraries remained diverse, largely non-memorizing and compatible with predicted peptide-like structural features. Together, these results establish pathogen-context conditioning as a new paradigm for computational library reshaping in antimicrobial peptide generation.

14
Gene Program Negotiation Defines Cellular Identity in Single-Cell Transcriptomes

Sung, J.-Y.; Cheong, J.-H.

2026-07-09 bioinformatics 10.64898/2026.07.05.736629 medRxiv
Top 0.1%
18.2%
Show abstract

Single-cell transcriptomics has transformed the characterization of cellular heterogeneity by enabling systematic analysis of biological gene programs. However, existing computational approaches primarily quantify the activity of individual programs independently and therefore provide limited insight into how multiple simultaneously active programs collectively determine cellular identity. Here we present Gene Program Negotiation (GPN), a graph-based computational framework that models regulatory decision-making among concurrently active biological programs. GPN reconstructs cell-specific program interaction networks from local transcriptional neighborhoods and quantifies regulatory organization using the Gene Program Coherence Index (GPCI) together with measures of local regulatory conflict, program diversity, and dominance. These graph-derived properties enable the classification of individual cells into five regulatory decision states: Consensus, Competition, Negotiation, Dominance, and Low activity. Applying GPN to gastric cancer single-cell transcriptomes revealed that cells sharing the same dominant biological program frequently occupied distinct regulatory decision states, demonstrating that dominant program identity alone does not uniquely define cellular regulatory organization. Competition states consistently exhibited elevated local regulatory conflict and were preferentially enriched among transition-like cells, indicating that regulatory competition is closely associated with transcriptional plasticity. Independent validation using glioblastoma single-cell transcriptomes reproduced these regulatory patterns without modification of the computational framework, supporting the robustness and generalizability of the approach across biologically distinct malignancies. These findings establish regulatory negotiation as an additional layer of cellular organization beyond conventional gene-program activity analysis. By explicitly modeling interactions among simultaneously active biological programs, GPN provides a general computational framework for investigating regulatory coordination, cellular plasticity, and dynamic cell-state organization in single-cell transcriptomic data.

15
Synthesizing Mechanistic Hypotheses from Single-Cell Omics via Discretized Feature Attribution and Empirical Language Model Grounding

Chen, J.; Hong, Y.; Bermudez, A.; Hu, J.; Hsieh, C.-J.; Lin, N.

2026-07-10 systems biology 10.64898/2026.07.09.737344 medRxiv
Top 0.1%
17.9%
Show abstract

Single-cell multimodal omics offer unprecedented resolution of cellular networks, yet translating continuous computational attributions into structured, testable biological mechanisms remains a persistent bottleneck. To address this limitation, we introduce an analytical pipeline employing decision trees to discretize continuous neural network attributions into explicit regulatory thresholds. These boundaries then structurally constrain large language models, enabling them to integrate established literature with empirical data to synthesize context-specific hypotheses. Applying this continuous-to-discrete framework across sparse datasets yielded novel biological mechanisms. Specifically, the framework articulated a cytoskeletal gating hierarchy governing EGF-stimulated pathways, identified transcriptomic drivers of input resistance in cortical interneurons, and delineated translational logic predicting Ki-67 abundance within spatial transcriptomics. Retrospective benchmarking validated the capacity of the framework to autonomously reconstruct published regulatory logic. Supported by a locally deployable open-weight language model and a code-free interface, this approach establishes an auditable methodology to extract robust experimental hypotheses from high-dimensional single-cell data.

16
Coding agents author interpretable single-cell embedding models from the literature

Brunn, N.; Krissmer, S. M.; Frosch, M.; Frick, M.; Prinz, M.; Binder, H.

2026-07-09 bioinformatics 10.64898/2026.07.07.737048 medRxiv
Top 0.1%
16.7%
Show abstract

The single-cell literature catalogs cell states as validated marker-gene programs - a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction. Here we show coding agents can author single-cell embedding models directly from the literature. Given a scenario that focuses this literature lens on a chosen biological subdomain, the agent edits a structured Python template, curating named, literature-cited gene programs and composing them into axes, without a gene-set database, training, or sight of the data. Across mouse and human tissues these zero-shot embeddings are competitive in biological quality with conventional, foundation-model, and program-informed baselines, batch-robust by construction and reproducible across runs, complementing data-driven embeddings. Because each dimension is a named, cited gene program, the embedding is interpretable and auditable, and its composable axes can be steered into a developmental tree.

17
FateLimit quantifies the prediction horizon of cell fate

Sung, J.-Y.; Cheong, J.-H.

2026-06-23 bioinformatics 10.64898/2026.06.22.733672 medRxiv
Top 0.1%
15.4%
Show abstract

Single-cell technologies have enabled increasingly detailed reconstruction of developmental trajectories, yet a fundamental question remains unresolved: when does future cellular identity become predictable from a cells current molecular state? Existing approaches infer lineage relationships, transition probabilities or future transcriptional dynamics, but do not directly quantify the emergence of fate predictability during cellular state transitions. Here we present FateLimit, an information-theoretic framework for measuring the temporal dynamics of cell-fate predictability from single-cell omics data. FateLimit combines probabilistic fate assignment, fate entropy and mutual information to quantify how information about future cellular outcomes is encoded in present molecular states. We introduce two quantitative descriptors: the Fate Information Half-Life (FIHL), which measures the characteristic timescale of fate-information dynamics, and the Prediction Horizon (PH), defined as the earliest developmental stage at which observed fate predictability exceeds the 95th percentile of a permutation-derived null distribution. We applied FateLimit across developmental, lineage-tracing and reprogramming systems, including pancreatic endocrinogenesis, CellTag reprogramming, human hematopoiesis and zebrafish embryogenesis. Across all datasets, FateLimit identified significant fate information and reproducible prediction horizons that were robust to cell-state representation, lineage structure and biological context. Comparative analysis revealed that prediction horizons differ substantially among cellular lineages, indicating that distinct developmental programs acquire predictive information at different rates. FateLimit establishes a general framework for quantifying the predictability of future cellular identity from present molecular states. By transforming developmental trajectories into predictability landscapes, FateLimit enables systematic comparison of commitment dynamics across biological systems and establishes prediction horizons as a quantitative measure of cell-fate determination.

18
Tokenizing single-cell transcriptomes as a native language for large language models

Xiao, C.; Ding, Y.; Bian, H.; Chen, Y.; Wei, L.; Zhang, X.

2026-07-11 bioinformatics 10.1101/2025.10.22.684047 medRxiv
Top 0.1%
15.1%
Show abstract

Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we propose CellTok, a tokenized single-cell language modeling approach that converts transcriptomic profiles into compact cellular token sequences and incorporates them into the vocabulary of a pretrained LLM. By representing cells as native tokens, CellTok enables cellular measurements, textual instructions, biological context, and multi-cell populations to be jointly processed within the same autoregressive modeling framework. Across diverse tasks, CellTok enable LLMs to recognize individual cells, interpret homogeneous and heterogeneous cell populations, infer disease-associated cellular states, predict cell-cell communication, model developmental trajectories, and generate cellular states. Moreover, prompt-based experiments show that providing appropriate biological context improves performance, indicating that CellTok can leverage LLM knowledge and contextual reasoning to support cellular data interpretation. These results demonstrate that single-cell transcriptomes can be transformed from a foreign molecular modality into a native language for LLMs, establishing a unified interface for modeling cells, populations, and biological knowledge in a shared token space.

19
A winding road to coexistence: Interdependence of niche and fitness differences in E. coli with targeted resource uptake gene deletions

McGuinness, B.; Guichard, F.; Weber, S. C.

2026-06-29 ecology 10.64898/2026.06.26.734880 medRxiv
Top 0.1%
14.8%
Show abstract

Resource competition theory typically assumes static traits and continuous supply of resources. Yet microbial communities often experience feast-famine cycles and rapid trait change. To investigate coexistence under these nonequilibrium conditions, we integrate modern coexistence theory (MCT) with a genome-scale metabolic model that explicitly links resource use (traits) to metabolic fluxes and growth. MCT partitions competitive interactions into niche and fitness differences, to predict when trait-driven departures from neutrality result in coexistence or exclusion. Using dynamic flux balance analysis, we define a function that maps trait-resource matching to niche and fitness differences between species in a two-species two-resource system. This mapping shows that niche and fitness differences are not independently tunable under resource competition: changes in transporter-mediated resource uptake and changes in resource concentration ratios generate constrained trajectories through coexistence space. Specifically, we show that the minimum niche difference required for coexistence increases linearly with the absolute difference in maximal growth rates on limiting resources, showing how limiting similarity between species can emerge from intracellular metabolic constraints. Furthermore, we find that in batch culture simulations, initial conditions (inoculum size, total resource concentration) determine the timescale of the transient growth phase, with niche differences saturating and fitness differences increasing as the timescale grows, thereby governing competition outcomes. Finally, we test these predictions experimentally using E. coli strains with targeted resource transporter knockouts under both equal and skewed resource concentrations. Our results confirm that transporter-mediated trait changes and resource concentration ratio modulation can be harnessed to engineer coexistence. Together, our work demonstrates that trait-resource matching imposes structured constraints on the joint evolution of niche and fitness differences, thereby shaping biodiversity maintenance in microbial communities under nonequilibrium conditions.

20
Machine Learning Gap-Fills Missing Transporter Kinetics in Biosystems Across Scales

Qiu, S.; Guo, Z.; Tu, W.; Zhuang, Y.; Wu, S.; Wang, G.

2026-07-07 systems biology 10.64898/2026.07.02.735998 medRxiv
Top 0.1%
12.9%
Show abstract

Understanding transporter kinetics is essential for deciphering metabolite exchanges in biosystems, particularly for cells subject to substrate gradients. Nevertheless, the prediction of transporter kinetic parameters, maximum rate per gram protein (Vmax) and Michaelis-Menten constant (Km), has not yet been tackled. Here, we developed the first compound-protein interaction machine learning model of transporter Vmax and Km, MMTKPred, which achieved R2=0.553, RMSE=1.155 mmol/hr/g Protein and R2=0.330, RMSE=0.935 mM for log10-scaled Vmax and Km prediction, respectively. Moreover, we demonstrated MMTKPred's predictive power across biosystem scales, from capturing transporter kinetics modulated by point mutations and substrate changes at the molecular level, to enabling substrate-sensitive metabolic modelling of non-model yeasts at the cellular level, and rationalizing inter-species substrate competition in co-cultures. Collectively, MMTKPred effectively models metabolite transport spanning from molecular to multi-species scales, thereby offering a computational tool for rational microbial cell factory optimization.